Papers by Ting-Hao Kenneth Huang

4 papers
LaMP-Cap: Personalized Figure Caption Generation With Multimodal Figure Profiles (2025.findings-emnlp)

Copied to clipboard

Challenge: Figure captions are crucial for helping readers understand and remember a figure’s key message.
Approach: They propose a dataset for personalized figure caption generation with multimodal figure profiles that provide inputs and profiles for each figure .
Outcome: The proposed dataset provides inputs and profiles for personalized figure caption generation with multimodal figure profiles.
From Selection to Generation: A Survey of LLM-based Active Learning (2025.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) have been used for selection and training of data for active learning.
Approach: They propose an intuitive taxonomy that categorizes LLM-based active learning techniques and discuss the transformative roles they can play in the active learning loop.
Outcome: The proposed model can generate entirely new data instances and provide more cost-effective annotations with fewer labeled data instances.
Using Contextually Aligned Online Reviews to Measure LLMs’ Performance Disparities Across Language Varieties (2025.naacl-short)

Copied to clipboard

Challenge: Of the world's 7,000 languages, sixty (60) million people speak British English, 23 million speak Taiwan Mandarin, and 10 million speak European Portuguese.
Approach: They propose a contextually aligned dataset that captures comments in different languages from real-world scenarios.
Outcome: The proposed approach shows that large language models underperform in Taiwan Mandarin in a sentiment analysis task.
"Newspaper Eat" Means "Not Tasty": A Taxonomy and Benchmark for Coded Language in Real-World Chinese Online Reviews (2026.acl-long)

Copied to clipboard

Challenge: Current language models handle coded language poorly, with limited real-world datasets and clear taxonomies.
Approach: They propose a taxonomy that captures common encoding strategies including phonetic, orthographic, and cross-lingual substitutions.
Outcome: The proposed model fails to detect or understand coded language in Chinese reviews . negative reviews can expose users to social pressure, retaliation, or reduced visibility .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations